Back

Artificial Intelligence in Medicine

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Artificial Intelligence in Medicine's content profile, based on 17 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Clinical Note Comparison and Data Retrieval Via Embedding Vectors: Model Selection, Metrics, and Convergence

Dahlberg, A. C. H.; Tapiola, O.; Luisto, R.; Puranen, T.; Sanmark, E.; Vartiainen, V.

2026-05-18 health informatics 10.64898/2026.05.12.26352832 medRxiv
Top 0.1%
9.7%
Show abstract

Background: Embedding models are an integral part of generative AI architectures, transforming text into embedding vectors that represent semantic content in numerical form. Despite their central role, their performance in clinical settings remains underexplored. We evaluate embedding models across two tasks: semantic difference detection in clinical texts, and data retrieval from patient records. Methods: Eight models were applied to synthetic discharge summaries in English, Finnish, and Swedish. Semantic sensitivity was assessed by introducing controlled perturbations (deletion, modification, and paraphrasing) at three levels of severity; cosine similarity, and L1 and Euclidean distances were computed between the vectors of the original and perturbed texts. Partial vectors were compared to explore dimensionality reduction. Two models with the biggest contrast in semantic difference detection were evaluated on retrieval of relevant information from real Finnish vascular surgery records. Results: Embedding vectors captured semantic differences in clinical text: content deletion and modification produced larger increases in vector distance than paraphrasing. On average, models detected the direction of semantic change correctly, but case-level performance varied considerably. Qwen3-Embedding-8B was the only model with zero directional errors, while multilingual-E5-large erred in 13.8% of cases. In data retrieval, Qwen3-Embedding-8B again outperformed multilingual-E5-large, though the margin was narrower: sufficiency scores were 3.25 vs. 3.17 out of 5 for the first query and 2.25 vs. 1.15 out of 5 for the second query. For some models, as few as 0.6-1.2% of dimensions sufficed to replicate full-vector accuracy; principal component analysis and coordinate-level analysis did not account for this finding. Conclusions: Our results show that the choice of embedding model is important: performance differences between models can be large enough to determine whether clinically relevant information reaches the end user, and model weaknesses can be both task-specific and context-dependent.

2
To RAG, or Not to RAG? A Comparative Evaluation of Retrieval-Augmented Generation for ICD Coding of German Tumor Diagnoses

Alickovic, F.; Lenz, S.; Ustjanzew, A.; Ortiz Rosario, L.; Vollmar, G. M.; Kindler, T.; Panholzer, T.

2026-06-03 health informatics 10.64898/2026.05.27.26353695 medRxiv
Top 0.1%
7.8%
Show abstract

Introduction Coding tumor diagnoses from free-text clinical documentation currently requires substantial manual effort. Promising approaches for automating this process include large language mod-els (LLMs), embedding models, and retrieval-augmented generation (RAG). While previous studies often focus on a single method, we directly compare these approaches on a real-world dataset of tumor diagnosis descriptions to assess their strengths and limitations. Methods We evaluated nine different embedding models using similarity search and embedding-based classification, as well as LLM-based coding, with and without RAG, on a real-world dataset of 2,024 unique German tumor diagnosis descriptions labeled with ICD-10 and ICD-O topography codes. The retrieval knowledge base was constructed exclusively from stand-ardized Alpha-ID, ICD-10-GM, and ICD-O-3 classifications. Performance was assessed for exact (full-code) and partial (three-character) code prediction. For RAG, we evaluated base and fine-tuned versions of Llama 3.1 8B and Llama 3.3 70B. Results Qwen3-Embedding-8B, the largest embedding model, yielded the best results. It achieved 47.8% exact-match and 72.1% partial-match accuracy for ICD-10 coding with classification, and 42.7% exact-match and 73.5% partial-match accuracy for ICD-O coding with similarity search. The other embedding models, including medically specialized ones, showed varied but lower performance. RAG improved base LLM perfor-mance and outperformed embedding-based approaches on partial-match accura-cy (80.6% partial-match accuracy for ICD-10 and 75.0% for ICD-O with Llama 3.3 70B), but not on exact-match accuracy. Conclusion A direct comparison with embedding-based approaches is essential to determine whether the additional effort of RAG is justified. The strong variation in performance also highlights the importance of model selection. Further advances in embedding-based methods, potential-ly supported by larger and more diverse training data, may offer a promising direction for future work.

3
Semantic Embeddings and the Peripheral Transcriptome in Ischemic Stroke: Connecting Molecular Signatures to NANDA-I Diagnoses

Santos, R. d. P.; Tinoco Patricio, A. d. O.; Gama, P. H.; Freitas, L. M. D.; Ribeiro, K. R.

2026-06-15 health informatics 10.64898/2026.06.11.26355453 medRxiv
Top 0.1%
7.4%
Show abstract

Objective: To construct and evaluate, in an exploratory manner, a pathophysiologic rationale link- ing biological pathways derived from the peripheral transcriptome in ischemic stroke (IS) to nursing diagnoses in the NANDA-I 2024-2026 taxonomy, while emphasizing that this association is not di- rect, deterministic, or automatically inferable from textual similarity with large language models (LLMs). Methods: A computational study was conducted using public secondary data from the Gene Ex- pression Omnibus series GSE16561, which includes 63 peripheral blood samples: 39 from indi- viduals with IS and 24 from healthy controls. The pipeline integrated transcriptomic analysis and functional enrichment, semantic mapping through ClinicalBERT embeddings, and mechanistic and clinical-conceptual judgment using Claude Sonnet 4.6 as a judge. The judgment stage was treated as the central interpretive layer, designed to mediate the transcriptome, pathophysiology, functional manifestation, and NANDA-I diagnosis. Results: The analysis identified a bimodal transcriptomic pattern, with activation of pathways re- lated to innate immunity and suppression of pathways related to adaptive immunity. Semantic map- ping generated 158 pathway-diagnosis pairs. The Spearman correlation between cosine similarity and the mechanistic score was negative and statistically significant (rho = -0.243; p = 2.09e-03), but weak in magnitude. This effect size indicates that semantic similarity explained less than 6% of the variance in mechanistic plausibility, reinforcing the insufficiency of embeddings as a stand- alone criterion. Of the 158 pairs, 14 were classified as high concordance, 8 as moderate, and 136 as divergent. Conclusion: The main value of this study lies in demonstrating that translating biological pathways into nursing diagnoses requires pathophysiologic, functional, and clinical-conceptual mediation. The prioritized pairs represent mechanistically plausible hypotheses for future research, without implying causality, direct clinical confirmation, or immediate care recommendations.

4
PhenoXtract: combining Large Language Model and Knowledge Graph embedding to extract phenotypes from clinical descriptions

Berardelli, S.; BRIERE, G.; Loire, B.; De Paoli, F.; Gazzo, A. M.; Limongelli, I.; Magni, P.; Zucca, S.; Baudot, A.

2026-06-26 genomics 10.64898/2026.06.22.733382 medRxiv
Top 0.1%
5.5%
Show abstract

Motivation: Standardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. Results: Evaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale.

5
Shortkit-ML: A Unified Multi-Perspective Framework for Detecting Shortcut Learning in Medical Imaging Embeddings

Cajas, S.; Marzullo, A.; Kapadia, S.; Santos, F.; Ocampo Osorio, F.; Kong, Q.; Quarta, A.; Kuo, P.-C.; Patel, M.; Rojas Sillery, R. I.; Celi, L. A.

2026-04-30 health informatics 10.64898/2026.04.29.26352053 medRxiv
Top 0.1%
4.7%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWShortcut learning poses a significant challenge in clinical artificial intelligence, as models may rely on spurious signals rather than clinically relevant features, leading to biased predictions and poor generalization. Existing detection methods are fragmented and lack systematic evaluation across datasets and model architectures. To address this issue, we propose ShortKit-ML, an open-source Python framework for unified shortcut analysis in embedding spaces. The framework integrates over 20 detection methods and six mitigation strategies within a modular pipeline, encompassing embedding analysis, fairness metrics, training dynamics, causal methods, explainability, and representation analysis. We evaluate the framework on chest X-ray datasets (CheXpert and MIMIC-CXR), synthetic benchmarks, and an out-of-domain dataset (CelebA). Experimental results demonstrate that multi-method auditing provides more stable and interpretable evidence than individual methods, while detector disagreement reveals meaningful representational differences. The proposed framework offers automated reporting, interactive visualization, and is available as a pip-installable package. The source code and documentation are publicly available at https://github.com/criticaldata/ShortKit-ML and https://criticaldata.github.io/ShortKit-ML/.

6
A multi-view similarity network fusion framework for syndrome discovery from aggregated health records

Gomes Ferreira, A. P.; Anzel, A.; Tavares Veras Florentino, P.; Pereira Ramos, P. I.; Barral-Netto, M.; Marcilio, I.; Hattab, G.

2026-07-02 health informatics 10.64898/2026.06.30.26356975 medRxiv
Top 0.1%
4.3%
Show abstract

Syndrome discovery, the identification of clinically meaningful groupings of signs and symptoms, is a foundational but labor-intensive task in syndromic surveillance, and the COVID-19 pandemic exposed the rigidity of expert-curated definitions in the face of novel threats. Unsupervised, data-driven methods are well-suited to this problem but remain underused. We propose an unsupervised framework based on Similarity Network Fusion (SNF) that operates on only five variables: diagnosis code, sex, age group, epidemiological week and year, and encounter count. Each diagnosis code was represented through three complementary views corresponding to the fundamental questions of syndromic surveillance: what condition is recorded (clinical, via SapBERT embeddings), who is affected (demographic, via chi-square distances), and when it occurs (temporal, via Move-Split-Merge). The fused affinity matrix is partitioned by spectral clustering and exported directly in the Open Syndrome Definition (OSD) format for downstream integration. To our knowledge, this is the first framework of SNF applied to the task of syndrome discovery. We use the framework in 72.9 million primary care encounters across ten Brazilian municipalities. Treating each city as an experiment with no shared training signal yields 47 candidate syndromes, 72% of which are rated fully valid by an expert blind to the procedure. By requiring no predefined targets, the framework discovers candidate syndromes at scale, including ones never explicitly sought, and emits them in a deployable format, shortening the path from emerging signal to usable definition.

7
Automated Interpretation of EEG Reports Using a Large Language Model with Structured Confidence Outputs

Tian, W.; Bergner, S.; Moiseev, A.; Popowich, F.; Medvedev, G.; Richardson, M. P.; Rodionov, R.; Xi, P.; Doesburg, S. M.; Ribary, U.; Winston, J. S.; Vakorin, V. A.

2026-07-10 health informatics 10.64898/2026.07.07.26357190 medRxiv
Top 0.1%
4.3%
Show abstract

Background: Free-text EEG reports typically lack structure, hindering scalable analysis. We evaluate a large language model (LLM) pipeline to extract structured diagnostic labels and confidence levels from these reports. Methods: We developed a hierarchical annotation schema to classify EEG reports for four specific abnormality types using a four-point confidence scale. To establish ground truth, two certified EEG technicians annotated a diverse dataset of reports authored by neurologists with distinct writing styles. We then implemented a grammar-constrained Mistral-7B pipeline, iteratively prompt-tuned on a development set to mirror these expert annotations. The pipeline's effectiveness was evaluated against the human expert benchmark using core agreement (diagnostic accuracy) and certainty-adjusted agreement (confidence alignment), with classical NLP models serving as a secondary baseline. Results: Mistral-7B significantly outperformed baselines, achieving 96% accuracy for overall abnormality detection, approaching the human benchmark of 98%. Crucially, the model successfully identified rare epileptiform abnormalities where traditional models failed and generalized robustly across distinct reporting styles. While diagnostic accuracy was high, a performance gap persisted in certainty-adjusted agreement, indicating that accurately modeling nuanced clinical confidence remains a challenge. Conclusion: LLMs can effectively automate the extraction of structured diagnostic information from EEG reports with near-human accuracy and strong generalization. While confidence calibration requires further refinement, the combination of accurate classification and explainability makes this pipeline a promising tool for standardizing clinical data at scale. Keywords: Routine Clinical Electroencephalography; Large Language Models; Clinical NLP; Confidence Assessment; Explainable AI; Neurophysiological Evaluation

8
Assessment of Zero-Shot Large Language Model (LLM) Assisted Clinical Trial Matching Processes: A Metastatic Cancer Use Case

Weng, Y.; Yalamaddi, H.; Fu, D.; Mishra, A.; Bunning, B. J.; Martin, A. B.; Hope, J.; Charu, V.; Kurian, A.; Desai, M.

2026-07-10 oncology 10.64898/2026.07.06.26354647 medRxiv
Top 0.1%
4.3%
Show abstract

Introduction: For oncology patients with limited treatment options, clinical trials may be a critical lifesaving pathway. Identifying relevant trials, however, is a time-consuming and difficult task. Several patient-trial matching processes incorporating large language models (LLMs) have been proposed to alleviate the burden on patients and oncologists. We aim to explore the benefits and practical challenges of zero-shot LLM-assisted trial matching processes by analyzing the results for a single pancreatic cancer patient. Materials and Methods: The results of a simple zero-shot LLM-assisted clinical trial matching process for our patient were compared to those of a "human benchmark," which was developed manually by two of the authors interfacing directly with ClinicalTrials.gov. Performance metrics -- sensitivity, specificity, precision, and accuracy -- were calculated. In addition, a qualitative content analysis (QCA) of LLM reasoning text was done to identify patterns in "errors," which we define as a human-LLM discrepancy in final patient eligibility. Implications and severity of errors are discussed. Results: The zero-shot LLM-assisted process returned potential trials with a sensitivity, specificity, and precision of 81.1%, 89.3%, and 86.5% respectively compared to the human benchmark. Qualitative error analyses revealed that about 73% of errors could potentially be alleviated with improved prompting and information access. Overall performance seemed comparable to that of human reviewers. Conclusion: The results from this preliminary real-world case study provide additional evidence to the literature in support of the integration of LLMs in clinical trial matching to provide benefit to patients with metastatic cancer with limited options.

9
Cross-Model Variability in Large Language Model Triage Behavior for Potential Stroke Symptoms

Dworkis, D. A.; Stenstrom, J.; Sen, A.; Lucarelli, R. T.

2026-05-25 emergency medicine 10.64898/2026.05.22.26353904 medRxiv
Top 0.1%
3.4%
Show abstract

Background: Stroke is a time-sensitive neurological emergency in which early EMS activation and presentation to definitive care are cornerstones of effective therapy. Large language models (LLMs) are increasingly consulted by the public for medical advice, but the veracity of the guidance provided by commercially available models responding to potential stroke symptoms is not well understood. Methods: We performed a cross-model benchmarking study comparing the triage choices of three frontier LLMs (Claude Sonnet 4.6, GPT-4o, and Llama 3.3-70b-versatile) on first-person vignettes describing a unilateral arm symptom on waking, across 10 symptom descriptors, and two clinical phases (before and after a partially reassuring self-examination), with or without a clinical distractor (n=50 per condition). Results: Claude sought emergency care most often, Llama least, and GPT-4o in between, diverging most sharply in the post-examination phase where Claude called 911 in 100% of runs, Llama called for non-emergency help in 100%, and GPT-4o was symptom-dependent. A distractor shifted behavior away from emergency care in almost all conditions: calling 911 fell from 37.9% to 14.6% and waiting rose from 0% to 45.9% in the post-examination vignette. Responses were also sensitive to symptom word: weak, limp, heavy, and clumsy generated higher alarm, whereas numb, tingly, odd, strange, and weird generated less urgent responses. Conclusions: The increasing use of LLMs for medical advice has significant public health implications. Commercially available LLMs show significant model-to-model variability and framing sensitivity when confronted with potential stroke symptoms, including under-recognition of canonical CDC warning descriptors, underscoring the need for systematic benchmarking as these tools become de facto first points of contact for patients experiencing neurological emergencies.

10
Large language models and retrieval augmented generation for complex clinical codelists: evaluating performance and assessing failure modes

Matthewman, J.; Denaxas, S.; Langan, S.; Painter, J. L.; Bate, A.

2026-04-24 health informatics 10.64898/2026.04.23.26351098 medRxiv
Top 0.1%
3.3%
Show abstract

ObjectivesLarge language models (LLMs) have shown promise in creating clinical codelists for research purposes, a time-consuming task requiring expert domain knowledge. Here, we evaluate the performance and assess failure modes of a retrieval augmented generation (RAG) approach to creating clinical codelists for the large and complex medical terminology used by the Clinical Practice Research Datalink (CPRD). Materials & MethodsWe set up a RAG system using a database of word embeddings of the medical terminology that we created using a general-purpose word embedding model (gemini-embedding). We developed 7 reference codelists presenting different challenges and tagged required and optional codes. We ran 168 evaluations (7 codelists, 2 different database subsets, 4 models, 3 epochs each). Scoring was based on the omission of required codes, and inclusion of irrelevant codes. We used model-grading (i.e., grading by another LLM with the reference codelists provided as context) to evaluate the output codelists (a score of 0% being all incorrect and 100% being all correct). ResultsWe saw varying accuracy across models and codelists, with Gemini 3 Pro (Score 43%) generally performing better than Claude Sonnet 4.6 (36%), Gemini 3 Flash, and OpenAI GPT 5.2 performing worst (14%). Models performed better with shorter target codelists (e.g., Eosinophilic esophagitis with four codes, and Hidradenitis suppurativa with 14 codes). For example, all models consistently failed to produce a complete Wrist fracture codelist (with 214 required codes). We further present evaluation summaries, and failure mode evaluations produced by parsing LLM chat logs. DiscussionBesides demonstrating that a single-shot RAG approach is currently not suitable for codelist generation, we demonstrate failure modes including hallucinations, retrieval failures and generation failures where retrieved codes are not used. ConclusionsOur findings suggest that while RAG systems using current frontier LLMs may create correct clinical codelists in some cases, they still struggle with large and complex terminologies and codelists with a large number of codes. The failure mode we highlight can inform the creation of future workflows to avoid failures.

11
Outcome Prediction Models for Critically Ill Patients Using Small Routine Laboratory Datasets

Cao, X.; Hou, J.; Wei, X.; Wang, Q.

2026-04-27 emergency medicine 10.64898/2026.04.26.26351758 medRxiv
Top 0.1%
3.3%
Show abstract

We present a suite of foundational, outcome prediction models for critically ill patients, developed using readily available, routine blood tests and advanced machine learning techniques. The input data of the models includes complete blood counts (CBCs), metabolic panels, and additional biomarkers that assess liver and kidney function, coagulation status, and cardiac injury. The output yields the predicted outcome at a given future horizon. For diagnoses, the length of the future horizon is set to zero while it is set to a fixed time interval for prognoses. The training dataset in this study comprises clinical data from 332 ICU patients, augmented with 200 synthetic samples generated via a conditional diffusion model. Generative machine learning-based data imputation and augmentation approaches yielded modest gains in predictive accuracy. However, substantial performance improvements were achieved through additional methods, including dimensionality and order reduction, SHAP-based feature importance analysis, and a novel time-series-to-image encoding strategy that enables the use of image-based classifiers for temporal clinical data. Principal component analysis-based order reduction produced measurable gains in outcome prediction, while the time-series-to-image encoding proved particularly effective in mitigating small-data limitations common in clinical research. Across all evaluation metrics--accuracy, precision, recall, F1 score, and AUROC--the prognostic models achieved performance exceeding 85%, with some models attaining AUROC scores above 90%. We innovated a new model-ensemble approach to optimize the predictive outcome. This ensemble modeling approach improves the overal prediction, pushing all assessment metrics over 90%. This work establishes a robust and interpretable AI-enabled diagnostic and prognostic toolkit for outcome predictions in critically ill patients and demonstrates a scalable workflow for developing high-performing models from sparse healthcare datasets. The proposed framework is readily deployable in ICU environments with routine blood testing capabilities and serves as a foundation for future integration into digital twin systems for critical care.

12
Hybrid Neural--Bayesian Belief Network Framework for Uncertainty-Aware Multimodal GBM Prediction

Jayme, A.; Heuveline, V.

2026-05-13 health informatics 10.64898/2026.05.10.26352710 medRxiv
Top 0.1%
3.3%
Show abstract

Background and ObjectiveGlioblastoma outcome prediction remains difficult because clinically relevant signals are distributed across heterogeneous imaging and genomic modalities, cohorts are small, and conventional neural predictors do not quantify their own uncertainty. This study evaluates a hybrid neural-Bayesian belief network framework for uncertainty-aware multimodal glioblastoma prediction and examines how modality selection, model family, and structure-aware regularization affect predictive performance and confidence quality. MethodsThe framework was evaluated on the TCGA-GBM radiogenomic cohort using four input modalities (T1Gd, FLAIR, mRNA, and CNA), five model families, five structural-weight settings, and 15 view subsets. A secondary benchmark on the UCI Human Activity Recognition dataset was included to assess whether observed limitations were specific to the glioblastoma setting. ResultsCNA features consistently reduced performance in most multimodal settings, and selective fusion excluding CNA outperformed both the full four-view baseline and imaging-only alternatives. Model families showed clear differences in uncertainty behaviour: non-Bayesian families achieved the strongest predictive accuracy, whereas the Bayesian family achieved the lowest calibration error over a narrower confidence range. Bayesian belief network regularization produced consistent directional improvements without supporting reliable structure-discovery claims, as learned graph structures were not reproducible across folds. On the secondary bench-mark, the same framework achieved much higher predictive performance, indicating that the glioblastoma performance ceiling primarily reflects data limitations rather than an architectural constraint. ConclusionsIn small-sample radiogenomic prediction, modality choice is at least as important as model choice, and uncertainty quality differs substantially across uncertainty-aware model families. The proposed framework provides a practical basis for comparing accuracy, calibration, modality selection, and structure-aware regularization in multimodal biomedical prediction.

13
NEXIM: A Nash Equilibrium-Based Framework for Stable Explainable AI in Medical Applications

Upadhyaya, D. P.; Sahoo, S. S.; Prantzalos, K.; Golnari, P.

2026-07-06 health informatics 10.64898/2026.06.25.26356568 medRxiv
Top 0.2%
3.2%
Show abstract

Reliable explanations are important for trustworthy medical applications of artificial intelligence (AI), but attribution-based explanations can vary across model randomization and small analytic changes. We present NEXIM (Nash Equilibrium-based Explainability and Interpretability Model), implemented here as an accuracy-constrained, equilibrium-inspired model-selection framework that jointly evaluates held-out prediction error, explanation stability, and cross-model connectivity. The implementation evaluated ten GradientBoostingRegressor models per prediction horizon, differing only by random seed (0-9), using a fixed 75/25 patient split. Kernel SHAP attribution vectors were compared using Spearman rank correlation, and graph connectivity summarized whether each model belonged to a dense explanation-similarity region. Candidate models within 0.02 Montreal Cognitive Assessment points of the best root mean squared error (RMSE) were ranked using a multiplicative Explanation Equilibrium Score. In longitudinal Parkinson's Progression Markers Initiative data, NEXIM selected the RMSE-optimal model at the one- and three-year horizons. At the two-year horizon, it selected Model 4 rather than the RMSE-only Model 8, increasing scaled stability from 0.8757 to 0.8847 and normalized graph connectivity from 0.889 to 1.000 while increasing RMSE by only 0.0014. The two models retained the same top-20 feature set but differed modestly in feature order, illustrating that NEXIM primarily acted as a reproducibility screen rather than identifying clinically contradictory explanations. Stability and consensus are treated as reproducibility criteria, not evidence of causal faithfulness, clinical usefulness, or improved patient outcomes. NEXIM may therefore serve as a governance checkpoint for model refresh and documentation, but external validation, stronger model-family baselines, and prospective clinical evaluation remain necessary.

14
Fine-Tuned Large Language Models for Detecting Social Isolation from Unstructured Clinical Notes

Chinthala, L. K.; Lemon, C.; Shaban-Nejad, A.; Farage, G.; Davis, R. L.; Xu, H.; Madlock-Brown, C.

2026-07-07 health informatics 10.64898/2026.07.05.26357334 medRxiv
Top 0.2%
3.2%
Show abstract

Objectives: This study aimed to leverage FLAN-T5-Large, BERT, RoBERTa, and Gemma-2-2B, with fine-tuning, to identify instances of social isolation and social support within unstructured clinical notes. Materials and Methods: Annotated clinical note spans containing social context cues were used to fine-tune each model. Performance was evaluated using Accuracy, Precision, Recall, and Macro-F1 score. A structured prompt was used to instruct the model to perform classification task and mitigate overgeneralization. Performance comparisons across the models assessed sensitivity, robustness, and false positive reduction. Results: FLAN-T5-Large achieved highest performance, with Macro-F1 of 0.92{+/-}0.04, demonstrating balanced results across classes: social isolation (F1 = 0.91{+/-}0.03), no social isolation (F1 = 0.94{+/-}0.05), and social support (F1 = 0.90{+/-}0.04). Gemma-2-2B produced comparable results, with Macro-F1 score of 0.89{+/-}0.10. BERT and RoBERTa achieved lower Macro-F1 scores of 0.77{+/-}0.17 and 0.80{+/-}0.21 respectively, with variability across categories. Discussion: A major contribution of this work is precise identification of multiple concepts related to social connectedness. By integrating annotated examples of both true and false positives, including negations and contextually ambiguous terms, the model better distinguished relevant social context cues from noise. Training on both social isolation and support provided a dual framework for comparative analyses and patient stratification. Conclusion: Transformer-based NLP models, particularly FLAN-T5-Large, demonstrated potential for identifying social isolation and social support in clinical text. These findings support the use of generative AI techniques to enhance detection of social isolation from EHRs, advancing context-aware healthcare analytics.

15
A Supervised Learning Framework for Stroke Hospitalization Factors Selection Using the Lasso-MIDAS Model

Li, Q.; Wang, L.

2026-05-20 cardiovascular medicine 10.64898/2026.05.15.26353365 medRxiv
Top 0.2%
2.9%
Show abstract

Stroke, as an acute cerebrovascular disease with significant public health implications, is influenced by a complex interplay of meteorological conditions, air quality, and socioeconomic factors. However, the inherent challenges of mixed-frequency data from diverse sources and high-dimensional variable spaces limit the effectiveness of traditional regression models. This study develops a Lasso-MIDAS model framework to identify the key multidimensional drivers of stroke admissions. Using this approach, 21 candidate variables encompassing meteorological, environmental, and economic indicators were screened. The empirical results identified 11 core influencing factors. In the meteorological and environmental dimensions, Wind Speed, Carbon Monoxide (CO), and Sulfur Dioxide (SO2) were identified as significant positive drivers, with Temperature Difference also positively correlating with admission risks. Conversely, Nitrogen Dioxide (NO2) exhibited a negative correlation, potentially reflecting behavioral adaptation and exposure reduction during peak pollution periods. In the socioeconomic dimension, the Consumer Price Index (CPI) for Food, Tobacco, and Alcohol emerged as a major risk factor, highlighting the impact of living cost pressures on public health. The findings demonstrate the superiority of the Lasso-MIDAS model in handling large-scale healthcare data. It effectively addresses the frequency mismatch problem while enhancing the robustness of causal identification through variable shrinkage. These conclusions provide a scientific basis for health authorities to establish early warning systems and optimize public health policy interventions.

16
Generation and Evaluation of Realistic Synthetic Clinical Progress Notes for Prostate Cancer using Large Language Models.

Rey-Blanes, A.; Veredas-Morente, J.; Vivas-Vargas, E.; Gil-Garcia, F.; Moreno-Barea, F. J.; Veredas, F. J.

2026-05-28 health informatics 10.64898/2026.05.25.26354027 medRxiv
Top 0.2%
2.8%
Show abstract

Background and Objective: Access to real-world electronic health records (EHRs) remains limited by privacy, governance and annotation constraints, hindering the development of clinical natural language processing models. Realistic synthetic progress notes may provide EHR-like corpora that preserve clinically rigorous information on diagnoses, treatments, symptoms, imaging, laboratory findings and therapeutic trajectories without relying directly on sensitive patient records. This study evaluates whether large language models (LLMs) can generate realistic Spanish prostate cancer progress notes from published case reports, preserving clinical content, temporality and hospital-style conventions.

17
Performance of Large Language Models as a Tool for Primary Care Consultations: Evaluation Study

Pascual, N.; Fernandez-Pichel, M.; Losada, D. E.; Garcia-Orosa, B.; Gude, F.; Costa Lathan, C.; Sueiro Justel, J.; Gomez Fontenla, A.; Lastra Perez, M.; Alonso Garcia,, F.

2026-05-04 health informatics 10.64898/2026.04.29.26352082 medRxiv
Top 0.2%
2.7%
Show abstract

Since the release of the first ChatGPT model in 2022, large language models (LLMs) have evolved significantly, and an increasing number of users now turn to these generative information systems for inquiries as sensitive and consequential as those related to health. The primary objective is to identify the main strengths and weaknesses of generative AI systems when responding to information needs as critical as those arising in the health domain. The study was structured using a question-answer format, in which each question corresponded to a user query and each answer represented the output generated by a model in response. The study employed a human evaluation framework involving two distinct panels of clinical experts from different specialties. The evaluation criteria encompassed three dimensions: adherence to medical consensus; presence or absence of inappropriate or incorrect information; and the potential to cause harm to users. GPT-4o mini, Llama 3, and MedLlama 3 were selected as three representative systems for the experiments. This study presents a detailed analysis of the performance of widely used contemporary large language models in addressing common health-related queries posed by online users. The results reinforce the potential of LLMs as tools for online health information seeking among non-expert users. However, the performance limitations identified underscore the need for further studies to monitor the future development of these models. Among them, performance issues have been identified in areas where users may be more vulnerable, leading to the retrieval of clinically incorrect information, particularly in matters relating to rare diseases. Furthermore, it has been noted that these models can become trapped in obsolete medical knowledge due to continuous scientific progress.

18
DentaCoPilot: An LLM-Augmented Next-Procedure Recommender for General Dentistry, Designed for Dentist Augmentation

Rodrigues, C. C.; Rebello, S. D.

2026-05-08 dentistry and oral medicine 10.64898/2026.05.07.26352635 medRxiv
Top 0.2%
2.6%
Show abstract

BackgroundCommercial dental artificial intelligence in 2026 is over-whelmingly diagnostic: caries, calculus, periapical, and bone-level detection on radiographs. The clinically harder question that follows every diagno-sis -- given a patients chart and most recent procedure, what should the dentist do next -- remains unsolved at general-dentistry scale. The closest published system, MultiTP (Chen et al., 2024), is a CNN-RNN restricted to partial-edentulism cases and provides neither calibrated uncertainty, structured rationale, nor an evaluation that treats the model as decision support rather than as an autonomous classifier. MethodsWe introduce DentaCoPilot, a recommender that, given a structured chart, returns (i) a calibrated top-K probability distribution over Current Dental Terminology (CDT) codes for the next procedure, (ii) a verbalised confidence label, (iii) an explicit abstain flag when context is insufficient, and (iv) a chartgrounded rationale. We compare four classical baselines (frequency bigram, TF-IDF + logistic regression, XGBoost, MultiTP-style CNN-RNN) and six large-language-model (LLM) variants (Claude Haiku, Sonnet + chain-of-thought, Sonnet + retrieval, Opus + chain-of-thought, Sonnet + classical prior, Opus + classical prior) on a synthetic chart corpus of 500 patients (1,284 test examples). All LLM inference is routed through the local Anthropic Claude Code CLI; every call is logged for full audit. ResultsOn apples-to-apples evaluation, classical baselines reach 0.567 top-1 / 0.967 top-5; pure LLM variants trail at 0.267-0.467 top-1. Prompt-conditioning a Sonnet LLM on the classical baselines top-10 candidates (M5) closes the gap: top-5 rises from 0.733 (pure Sonnet + chain-of-thought) to 0.933, matching classical baselines, while preserving rationale and abstention. Increasing the LLM backbone from Sonnet to Opus does not improve accuracy with or without priming. Calibration via temperature scaling and coverage-risk analysis is reported for the baselines. ConclusionPrompt-conditioning a small LLM on a classical baselines top-K is the most cost-effective LLM design we tested for next-procedure recommendation, and the design preserves the augmentation features that distinguish the system from an autonomous classifier. A pre-registered clinician-in-the-loop evaluation at the KLE Vish-wanath Katti Institute of Dental Sciences (Belgaum, India) and a real-data evaluation on the multi-institutional BigMouth dental data repository are the next stage of work.

19
A Consensus-Driven Stacking Ensemble Framework for Interpretable Cardiovascular Risk Prediction and Clinical Deployment

Sozol, S. S.; Dev Nath, B. C.; Fahim, F. M. S.; Suzana, N. N.; Mirza, J. F.; Ahmmed, S.; Zohra, F.-T.; Zafr, A. H. A.; Uddin, M. N.; Mondal, M. R. H.; Hoque, A. S. M. L.

2026-05-26 health informatics 10.64898/2026.05.18.26352989 medRxiv
Top 0.2%
2.6%
Show abstract

Machine learning (ML) is being considered to help diagnose cardiovascular diseases (CVD). Still, challenges like inconsistent and limited datasets, limited infrastructure, and global inequalities lead to the need for a reliable and practicable ML solution. This paper presents an ML-driven framework for predicting CVD risk scores and classifying status. Several data preprocessing techniques, including multiple imputation by chained equations (MICE), outlier removal, are considered. In addition, hyperparameter tuning is performed with the GridSearchCV tuning technique. Moreover, a consensus-driven five-feature selection method is applied to identify optimal predictors. The dataset used in this study contains healthcare records related to future CVD risk scores, comprising 1,529 patient records with 22 features. The optimized stacked ensemble model is applied to the dataset and achieves a cross-validated coefficient of determination value of 98.13% for CVD risk score regression. Comparative evaluation with other ML models confirmed improved accuracy, efficiency, and interpretability. The explainable AI technique SHAP is applied to interpret predictions and highlight key risk factors. Moreover, a deployment-ready web platform with multi-role access has been developed that demonstrates clinical applicability. The proposed framework offers a reliable and interpretable tool for early detection of CVD and personalized risk assessment. In the future, this work can be extended to integrate longitudinal data, medical imaging, and deep learning to improve generalizability and strengthen real-world impact.

20
Rare-Class Collapse in ECG-Based Ventricular Tachycardia and Fibrillation Detection: A Systematic Benchmark of Class-Imbalance Mitigation from Reweighting to Cascade Classification

Tiruwa, K. R.

2026-06-29 cardiovascular medicine 10.64898/2026.06.26.26356694 medRxiv
Top 0.2%
2.5%
Show abstract

Ventricular tachycardia (VT) and ventricular fibrillation (VF) are the leading electrical causes of sudden cardiac death, but automated detection is limited by strong class imbalance, where lethal arrhythmias account for fewer than 22% of ECG segments. In this setting, standard classifiers can achieve high accuracy by predicting normal rhythm in most cases while missing many lethal events, a failure mode referred to as rare-class collapse. We evaluated six imbalance-handling approaches: naive logistic regression, inverse-frequency reweighting, label-distribution-aware margin loss (LDAM), cost-sensitive training, two-stage cascade classification, and anomaly detection on 15,614 ECG segments from three PhysioNet databases (VTaC, VFDB, CUDB), with an overall normal-to-lethal ratio of 3.6:1. All methods were assessed at a fixed operating point of 95% specificity using recall, area under the precision-recall curve (AUPRC), and missed-lethal-event rate (MLER). The naive model achieved 45.1% recall (MLER = 0.549), missing 564 of 1,027 lethal events despite 84.1% accuracy. The two-stage cascade performed best, with 65.2% recall, AUPRC of 0.821, and MLER of 0.348, reducing missed events by 37% and achieving the highest decision-curve net benefit. Per-source analysis showed near-complete VF detection (recall up to 0.975) but much lower VT detection (recall 0.183), suggesting a feature-space limitation due to spectral similarity between organized VT and rapid sinus rhythm. Overall, the results show that evaluation metrics strongly influence the visibility of rare-class failure, and that cascade-based methods outperform simpler reweighting approaches for detecting lethal arrhythmias.